Back

PLOS Digital Health

Public Library of Science (PLoS)

Preprints posted in the last 90 days, ranked by how well they match PLOS Digital Health's content profile, based on 106 papers previously published here. The average preprint has a 0.26% match score for this journal, so anything above that is already an above-average fit.

1
Assessing electronic health record potential for adaptive learning in multimorbidity care in Sub-Saharan Africa: a mixed-methods study of Zimbabwe's Impilo system

Dhodho, E.; Choga, K.; Mundoga, F.; Chimberengwa, P. T.; Gongora, R. T.; Webb, K.; Chinyanga, T. T.; Banda, F.; Masiye, K.; Midzi, N.; Mudavanhu, J.; Katsidzira, A.; Manyiyo, B.; Apollo, T.; Chimbetete, C.; Mhlanga, T.; Mangisi, P.; Gwanzura, C.; Tsvangirayi, S.; Dixon, J.; Nitsch, D.

2026-07-19 health informatics 10.64898/2026.07.16.26357920 medRxiv
Top 0.1%
61.1%
Show abstract

Electronic health records (EHR) are increasingly recognised as critical digital infrastructure for integrated, patient-centred care in the context of rising multimorbidity. In low-resource settings, national EHRs may also support locally driven learning to improve adaptive care across chronic conditions. However, there is limited empirical evidence on whether and how these systems enable learning within routine care in ways that inform broader system adaptation. We conducted a qualitative multi-method assessment of Impilo, Zimbabwe's national EHR, to examine its capacity to support learning for integrated multimorbidity care at primary care level, using HIV-hypertension as a tracer condition pair. Guided by Friedman's socio-technical infrastructure model as the analytical framework and Learning Health Systems (LHS) theory as the interpretive framework, data were drawn from documentary review, ethnographic observation, patient journey mapping, and interviews with frontline health workers and key stakeholders. Frontline learning for person-centred multimorbidity care was actively generated through interpretation of patient trajectories, experiential adjustment, and coordination across HIV and hypertension services using both the EHR and paper-based artefacts such as registers and patient booklets. However, this learning remained largely encounter-bound and weakly stabilised. Impilo did not routinely provide usable longitudinal patient views, practice-facing analytic tools, or institutionalised mechanisms for collective reflection required to support integrated multimorbidity care. Consequently, learning was largely confined to incremental adjustment within existing workflows, with limited capacity to inform broader changes to care pathways, routines, or system design. These findings suggest that the principal barrier to developing LHS is not the absence of data or frontline learning capacity, but the lack of socio-technical arrangements that enable learning to stabilise and inform system adaptation. Digitalisation alone is insufficient to support adaptive multimorbidity care. Co-production with frontline health workers may provide a pathway for aligning digital system design with routine care realities.

2
NigBench: A multilingual point-of-care medical query benchmarking study of large language models in Nigeria

Olatunji, T.; Aka, C.; Okocha, C.; Ayodele, E.; Orisakwe, J.; Adekunle, T.; Sanni, M.; Abiola, A.; Abdullahi, T.; Owopetu, O.; Afolaranmi, T.; Yougha, P. S.; Emmanuel-Fabula, M.; Menon, V.; Denniston, A.; Liu, X.; Williams, G.; Mateen, B. A.

2026-07-10 health informatics 10.64898/2026.07.05.26356776 medRxiv
Top 0.1%
60.2%
Show abstract

In this study, we introduce a novel benchmark comprising over 9,000 real-world, point-of-care, multilingual, and multimodal clinical question-answer pairs sourced from frontline health workers in Nigeria. Using the dataset, we compare local general practitioners to multiple leading open and closed LLMs. Our results reveal several critical insights into the suitability of LLMs as clinical decision support systems in low-resource contexts. The results confirm that performance varies widely by language and input modality (e.g., text vs speech): while models perform best on English text inputs, their accuracy drops significantly for local-language speech. Critically, it is possible to achieve substantial performance gains by transcribing and translating other languages into English before prompting an LLM-- an important insight for non-anglophone product developers. Finally, this benchmark highlights key limitations of SLMs in supporting frontline healthcare in low-resource settings and provides a clear opportunity to track improvements as novel solutions are developed.

3
What it takes to implement AI in Africa: health-system lessons from developing an ML-enabled maternal risk stratification algorithm in Tanzania

Hellar, A. M.; Lyatuu, I.; Kinyina, A.; Kulindwa, Y.; Ernest, E.; Mandali, H.; Mtani, C.; Athumani, H.; Phiri, F.; Massawe, P.; Sukari, O.; Sospeter, P.; Kapologwe, N.

2026-07-28 health systems and quality improvement 10.64898/2026.07.26.26358974 medRxiv
Top 0.1%
58.6%
Show abstract

Background Machine learning (ML) has growing potential to support early identification of high-risk pregnancies in resource-constrained settings. However, most studies focus on model development and predictive performance, with less attention to the health-system processes required to generate ML-ready data and translate risk information into clinical action. The Mlinde Mama Project in Tanzania combined Group Antenatal Care (G-ANC), digital maternal health systems, and development of an ML-enabled risk stratification model for hypertensive disorders of pregnancy (HDP). This study examined the health-system conditions shaping the pathway from routine care to actionable ML-enabled risk information. Methods We conducted a retrospective mixed-methods implementation analysis of this project that was implemented between 2022 and 2024 in Geita, Tanzania. The analysis triangulated endline evaluation findings, quantitative exit interviews with pregnant women, focus group discussions with women and healthcare workers, key informant interviews, project implementation records, routine data from Tanzania's Unified Community System (UCS), and documented ML development experience. Relevant quotations from the final evaluation report were systematically screened, selected, and coded. An abductive analysis combined inductively derived themes with the Non-adoption, Abandonment, Scale-up, Spread and Sustainability (NASSS) framework and a study-specific ML implementation pathway. Results Seven themes emerged across three phases of the implementation pathway. ML readiness began before the algorithm, with reliable clinical measurement and documentation; patient participation and task-sharing redistributed data generation. The paper-to-digital transition shaped which information became ML-ready data. The potential value of ML depended on workflow redesign rather than prediction alone. Technology created an efficiency paradox in which task-sharing reduced workload while staffing shortages and confirmatory work generated new burdens. Advanced analytics depended on basic infrastructure, while sustainability relied on teamwork, trust, user demand, and local ownership. Conclusions Our findings suggest that successful ML implementation in maternal health begins with health-system readiness before the algorithm itself. A critical paper-to-digital transition stage was a major determinant of data quality and ML readiness. ML-enabled maternal risk stratification should therefore be approached as a sociotechnical intervention spanning measurement, digitization, clinical confirmation, and follow-up. This implementation readiness is critical for the successful development of an ML-enabled risk stratification system. Crucially, our findings resonate with the key domains of the NASSS framework, underscoring how addressing multidimensional complexities, from technological design to organizational readiness, is vital for successful scale-up and long-term sustainability.

4
Beyond isolated cough events: AI-based tuberculosis screening through temporal analysis of cough sounds

Ma, N.; Mirheidari, B.; Brown, G. J.; Muyoyeta, M. M.; Sanjase, N.; Maimbolwa, M. M.; Chifwamba, S.; Muzazu, S.; Kagujje, M.

2026-07-13 infectious diseases 10.64898/2026.07.08.26357519 medRxiv
Top 0.1%
50.5%
Show abstract

Tuberculosis (TB) is a major global health challenge, with many cases remaining undiagnosed due to limited access to screening and diagnostic services. Artificial intelligence (AI) systems based on cough sound analysis offer a scalable and accessible approach to TB screening, but most previous studies have analysed isolated cough events, despite the possibility that diagnostically useful information is encoded in the temporal dynamics of cough episodes. We evaluated an AI-based screening framework using cough recordings collected under real-world clinical conditions from 500 participants in Zambia, including 201 individuals with bacteriologically confirmed TB, 150 symptomatic patients with other respiratory diseases, and 149 healthy controls. Using multiple pre-trained speech foundation models fine-tuned on cough sounds, we systematically investigated the influence of temporal context by varying the audio input window from 1 to 6 s, measured from the onset of each cough episode. Across all evaluated models, diagnostic performance consistently peaked with a 3 s input window, indicating that useful information extends beyond individual cough events and is encoded within the short-term temporal dynamics of cough episodes. The best audio-only model achieved an area under the receiver operating characteristic curve (AUROC) of 85.2% for distinguishing TB from all other participants and 80.1% for distinguishing TB from symptomatic non-TB respiratory disease. Incorporating demographic and clinical variables improved AUROC to 92.1% and 84.2%, respectively. Performance remained robust across recording devices, participants with HIV co-infection, and varying acoustic conditions. These findings demonstrate that preserving temporal context improves AI-based cough screening for TB and suggest that analysing cough episodes, rather than isolated cough events, may enhance diagnostic performance in real-world settings. More broadly, the results highlight the importance of temporal context in the design of future respiratory sound datasets and AI-based diagnostic systems.

5
A digital health approach for identifying polyendocrine metabolic ovarian syndrome using machine learning and body temperature

Awoniran, O. M.; Lawlor, D. A.; Gaunt, T. R.; Millard, L. A. C.

2026-07-27 health informatics 10.64898/2026.07.23.26358666 medRxiv
Top 0.1%
45.8%
Show abstract

Background Polyendocrine Metabolic Ovarian Syndrome (PMOS), formerly known as Polycystic Ovary Syndrome (PCOS), is a prevalent endocrine disorder with high rates of undiagnosed cases globally. Accessible screening tools are needed to facilitate appropriate management and earlier intervention. As PMOS is frequently characterised by oligo-anovulation, the absence of the characteristic rise in basal body temperature typically seen in ovulatory cycles may serve as a physiological marker for the condition. Objective This study aimed to assess the feasibility of using machine learning to identify individuals with PMOS from temperature data collected by a body-worn device. Methods We used data from 387 users of a vaginal temperature monitor (OvuSenseTM) who responded to a questionnaire. The sample was restricted to individuals with at least three cycles with sufficient temperature data and whose PMOS case/control status could be determined from questions about prior clinical consultation for infertility and conditions for which they take medications. We randomly sampled three menstrual cycles for each participant and derived a set of cycle-level and user-level temperature features. Cycle-level features included cycle length and measures describing the temperature rise indicative of ovulation (e.g. temperature rise, cycle day of temperature rise start). We also constructed a reference cycle representing the typical bi-phasic cycle pattern (created using cycles from those without known fertility conditions) and used this to derive features describing how much a participant's cycles differed from this reference. The cycle-level features were aggregated into user-level features by taking the minimum, maximum, median, and range of the cycle-level features across the three selected cycles for each participant. We used 5-fold nested cross-validation to evaluate the extent that PMOS could be predicted, at the cycle and user levels, using Logistic Regression (LR), Support Vector Machine (SVM), and Random Forest (RF). Results The average age of participants was 31.97 years (SD=4.58), with 49.6% having a self-reported PMOS diagnosis. The models demonstrated moderate discrimination, with cycle-level AUC-ROC scores ranging from 0.64 (SD=0.02) (LR) to 0.68 (SD=0.04) (RF), and user-level scores ranging from 0.65 (SD=0.07) (LR) to 0.70 (SD=0.04) (RF). All models were reasonably calibrated, though confidence intervals were wide (e.g. RF cycle-level: calibration slope = 0.83 (95% confidence interval [CI]: 0.68, 1.00), calibration intercept = 0.02 (95% CI: -0.11, 0.14); user-level: slope = 0.88 (95% CI: 0.69, 1.15), intercept = -0.01 (95% CI: -0.22, 0.16)). Conclusions This study demonstrates the potential of using body temperature from digital health devices to identify those with PMOS. Such a passive approach to identifying PMOS could help to identify undiagnosed PMOS in those who have not actively sought a diagnosis. Further research is needed to assess its predictive performance and acceptability in a general population using more widely used digital devices.

6
Human In the Loop Challenges for Quality Annotation of Pre-Cancer Lesions in Clinical Oral Images

Mandal, S.; Mendonca, P.; Gurushanth, K.; Thakur, H.; Birur, P.; Shetty, A.; Pal, D.

2026-07-04 dentistry and oral medicine 10.64898/2026.07.02.26355859 medRxiv
Top 0.1%
45.6%
Show abstract

Background: The hyperplasia and dysplasia stage (pre-cancer) offers a viable opportunity to reduce the incidence and mortality of oral cancer through early prevention. Smartphone-based Artificial Intelligence (AI) enabled screening of potentially malignant oral lesions offers a scalable solution for this in resource-constrained settings. However, developing accurate and explainable AI segmentation models require high-quality, pixel-level annotated data. This process that is prohibitively expensive, time-consuming, and prone to inter-observer subjectivity among clinical experts. Methods: We designed and empirically validated a Deep Learning-driven Human-in-the-Loop (HITL) framework to pixel-annotate a dataset of 3026 clinical oral images. Using an iterative pseudo-labeling pipeline, we evaluated the model's learning dynamics and performance evolution across five training cycles. We conducted controlled experiments to quantify the networks tolerance to intermediate level of label noise (unreviewed pseudo-labels) to resolve clinical subjectivity using pixel-wise Cohen's Kappa and the STAPLE consensus algorithm. Results: Iterative self-training produced sustained improvements in lesion detection and spatial localization. However, controlled experiments revealed that including even a modest fraction ({approx}10%) of unreviewed pseudo-labels led to a three-to-four-fold increase in training convergence instability and induced a conservative prediction bias that negatively impacted model recall. When measured against multi-expert ground truth, the model's performance converged with the inter-rater reliability ceiling ({kappa} {approx} 0.65), indicating that its predictions fell within the envelope of human agreement. Conclusions: Our findings emphasize that a final expert-driven quality assurance step remains absolutely essential to mitigate training instability, confirmation bias, and clinically unacceptable drops in recall caused by label noise. Overall, this work provides a scalable, empirically validated blueprint for building domain-specific medical imaging datasets in low-resource global health settings, where the dual challenges of annotation cost and inter-observer variability are most acute.

7
Digital Health Adoption, eHealth Literacy, and Trust in AI Among Generation Z University Students in Sri Lanka: An Empirical Study

Athukorala, S. C.

2026-07-28 health informatics 10.64898/2026.07.22.26358733 medRxiv
Top 0.1%
45.1%
Show abstract

Background: Digital health technologies, spanning mobile applications, telemedicine, and AI-driven platforms, are rapidly reshaping healthcare delivery globally. Although Generation Z university students are classified as digital natives, empirical data evaluating their eHealth literacy, technology acceptance, and specific trust barriers in developing South Asian nations like Sri Lanka remain scarce. Objective: This study evaluated eHealth literacy, technology acceptance, online health information-seeking behaviors, and adoption barriers among Gen Z undergraduates in Sri Lanka, focusing on the interplay between eHealth literacy, AI trust, and digital care preferences. Methods: A cross-sectional survey (N = 172) was conducted among Sri Lankan university undergraduates utilizing adapted, validated instruments: the eHealth Literacy Scale (eHEALS) and the Technology Acceptance Model (TAM). Statistical analysis included scale reliability validation (Cronbach's alpha), descriptive profiling, Chi-Square ({chi}{superscript 2}) contingency tests, Pearson correlations, and Multiple Linear OLS Regression models. Results: Participants demonstrated high overall eHealth literacy (Mean = 3.84 {+/-} 0.58) and strong endorsement of digital health utility (Mean = 3.99 {+/-} 0.59). Online health searches were reported by 86.6% of respondents. AI tools (e.g., ChatGPT, Gemini) emerged as the second most frequent source for health queries (57.6%), surpassing YouTube (44.2%) and social media (26.2%), with medical students showing significantly higher AI utilization ({chi}{superscript 2} = 8.70, p = .003). In multiple regression analysis, digital platform preference over physical clinic visits (R{superscript 2} = .352, p < .001) was significantly predicted by Perceived Ease of Use ({beta} = 0.371, p = .001) and Trust in AI Recommendations ({beta} = 0.370, p < .001), whereas face-to-face consultation preference (76.7%) and personal data privacy risks (50.0%) remained predominant adoption barriers. Conclusion: Gen Z students in Sri Lanka exhibit high digital health readiness and substantial reliance on AI-driven information seeking. However, institutional deployment must address privacy concerns and integrate hybrid clinical workflows to bridge the gap between high perceived utility and physical consultation preferences.

8
Malaria Pre-screening Technology Using Artificial Intelligence (AI)

Ibeto, O. O.; Nwoye, E. O.

2026-07-17 infectious diseases 10.64898/2026.07.15.26357432 medRxiv
Top 0.1%
45.1%
Show abstract

Malaria remains a severe health problem in endemic regions because people lack adequate diagnostic tools, leading to delayed medical care and elevated death rates. This research introduces a dual-mode artificial intelligence system that uses two complementary models to enhance malaria pre-screening and diagnosis. The patient-centered model uses multivariate logistic regression to analyze biosignals, including heart rate, body temperature, and oxygen saturation, collected through a wearable sensor prototype and a mobile interface for symptom analysis. The system enables patients to begin self-assessment to determine their level of need before scheduling a doctor's appointment. The clinician-centered model represents a customized convolutional neural network that uses annotated microscopy images of red blood cells to achieve 94.84% accuracy, 95.71% precision, 93.87% recall, 94.78% F1 score, and 0.84 Area Under Curve (AUC). The patient model achieved 94.6% accuracy and an AUC of 0.985 using a 70/30 train-test split. These systems work together to create a layered diagnostic system that can operate independently or together to detect malaria at an early stage, especially in areas with limited resources. The findings demonstrate that wearable biosignal data integration with image-based deep learning can produce dependable, scalable, and user-friendly systems for malaria pre-screening. Keywords - malaria diagnosis, artificial intelligence (AI), convolutional neural networks (CNN), wearable biosensors, multivariate logistic regression

9
Behavioural readiness, not demographics, predicts wearable adoption and digital medicine integration in a diverse multinational population: a cross-sectional study of 3,004 adults in Qatar

Zaghloul, H.; Arabi, B.; Al-Ani, M.; Abdullah, A.; El-Masri, R.; AboMuslim, O.; Al-Ahdab, F.; Rizwan, M. R. M.; Tag, Z.; Zaghlool, S.; Arayssi, T.

2026-07-20 health informatics 10.64898/2026.07.17.26358328 medRxiv
Top 0.1%
44.6%
Show abstract

Whether diverse populations outside Western settings are behaviourally ready to integrate wearable-derived data into clinical care remains poorly understood. This study examines sociotechnical determinants of wearable adoption and digital health data-sharing readiness in a large, highly diverse multinational population in Qatar, a rapidly digitising health ecosystem with advanced eHealth infrastructure. We conducted a cross-sectional community-based survey of 3,004 adults across Qatar, assessing wearable device use, behavioural engagement, and willingness to integrate wearable-generated data into healthcare workflows. Multivariable logistic regression identified independent predictors of wearable adoption. Wearable device use prevalence was 34.1%. Behavioural factors were the strongest independent predictors of adoption: daily exercisers had more than four times the odds of wearable use compared with rarely active participants, and willingness to share data with healthcare providers was independently associated with adoption after full adjustment. Notably, education level was not independently associated with wearable use, suggesting that behavioural readiness outweighs traditional socioeconomic indicators as a determinant of digital health engagement. Older age ([&ge;]56 years) and African ethnicity were associated with lower adoption odds, highlighting persistent digital inequities. These findings challenge the assumption that digital health equity is primarily an education or access problem, repositioning it as a behavioural engagement challenge. Health systems scaling remote monitoring programmes should prioritise identifying behaviourally engaged subpopulations rather than relying solely on demographic targeting. Targeted digital engagement strategies addressing older adults and underrepresented ethnic groups are essential for equitable implementation of digital medicine.

10
VideoCap: Enabling REDCap as a Tool for Crowdsourced Multimedia-Labeling Research

Alasaly, B.; Jang, K. J.; Zolensky, A.; Mopidevi, S.; Johnson, K. B.

2026-07-09 health informatics 10.64898/2026.06.26.26355995 medRxiv
Top 0.1%
39.6%
Show abstract

Objective: Machine learning systems that use video data require large, diverse, human-labeled datasets, but generating reliable annotations remains labor-intensive, difficult to scale, and often dependent on proprietary tools or small expert annotator pools. We present VideoCap, a secure and context-aware workflow that integrates REDCap with a Content Delivery Network (CDN), backend web service, and crowdsourcing platform to dynamically rotate embedded video segments within a single survey structure. Methods: VideoCap hosts segmented videos on a CDN with restricted downloads, stores segment metadata and session state for automated video selection, and generates REDCap survey URLs populated with embedded video parameters. We implemented the workflow through Amazon Mechanical Turk to annotate simulated patient-provider video segments using open-ended insights, structured scheme selections, and free-text responses. Results: We tested the workflow using 481 simulated patient-provider video segments over a 131-day deployment period. In the valid-only analytic subset, 358 unique video segments received 814 annotations, with an average of 2.27 annotations per segment. The workflow achieved an average annotation-time-to-video-duration ratio of 4.47:1, lower than contextual annotation-time estimates reported in prior multimedia annotation workflows. Conclusion: VideoCap provides a reproducible workflow for dynamic multimedia annotation using broadly accessible tools, demonstrating feasible survey delivery and efficient annotation.

11
PhysiCase: Development and dual-layer validation of synthetic cases for health professional education: A pilot study leveraging Generative AI

Komolafe, O. O.; Roberts, A. C.; Shelley, J.; Tawiah, A. K.

2026-06-09 rehabilitation medicine and physical therapy 10.64898/2026.06.07.26355114 medRxiv
Top 0.1%
39.0%
Show abstract

High-quality, domain-specific datasets are foundational to advancing educational tools and AI systems in healthcare, yet assembling case repositories from real-world clinical records faces substantial privacy, ethical, and licensing barriers. Synthetic data generation offers a compelling pathway forward, but educational cases require rigorous validation to ensure clinical plausibility and pedagogical utility. This pilot study introduces PhysiCase, a dual-layer validation pipeline for synthetic case generation and evaluates the feasibility of combining automated LLM-based screening with expert educator review. We generated 128 synthetic musculoskeletal(MSK) cases using four frontier large language models (GPT-4.1, GPT-4o, Google Gemini 2.5 Pro, and Llama 4 Scout) across 28 clinical conditions. Cases underwent automated quality screening using an "LLM-as-judge" framework (DeepEval) assessing prompt alignment, JSON correctness, answer relevance, bias, toxicity, and completeness. Ninety cases (70.3%) passed automated filtering and proceeded to expert evaluation by four MSK physiotherapy educators, who rated medical accuracy, realism, fidelity, relevance, and usability on 5-point Likert scales. GPT-4.1 demonstrated the highest automated pass rate (96\%) and strongest expert ratings (medical accuracy 4.10/5, usability 4.38/5), while Llama 4 Scout showed the lowest pass rate (33.3%) and expert ratings. Expert-evaluated cases achieved strong content validity indices for usability (97.5%), relevance (97.5%), and realism (95%), though medical accuracy showed greater variance (CVI 87.5%). Cross-layer correlation analysis revealed that automated completeness metrics moderately aligned with expert usability ratings , while answer relevance and prompt alignment showed weak or negative correlations with clinical correctness. Qualitative analysis identified three primary failure modes: reductive logic, biomechanical inconsistency, and administrative/contextual gaps. The dual-layer validation framework proved methodologically viable: automated screening efficiently reduced expert review burden, while human judgment remained indispensable for detecting subtle clinical reasoning failures. LLM-generated synthetic cases has the potential to meet practical educational needs for MSK physiotherapy, but expert validation is essential to safeguard clinical accuracy. These findings support a scalable division of labour for synthetic case development, with targeted improvements to prompting and automated reasoning checks needed to address identified "nuance gaps." The code for this paper is available on https://github.com/kwid-ai/PhysiCase

12
Expert-Guided Visual Correction for Characterizing Diagnostic Performance and Error Patterns of Multimodal Large Language Models Using Periodontal In-Service Examination Images

Dhaimade, P. A.; Henderson, R.

2026-08-27 dentistry and oral medicine 10.64898/2026.08.21.26360755 medRxiv
Top 0.1%
38.9%
Show abstract

Multimodal large language models (MLLMs) are increasingly applied to image-based clinical reasoning, yet their diagnostic reliability in periodontal image interpretation, and the underlying source of their errors, remain poorly characterized. This study evaluated six architecturally distinct MLLMs (Claude Sonnet 4.5, GPT-5.0, Gemini 2.5, GLM-4.6, Sonar, and Grok 4.1) using 50 image-based multiple-choice questions drawn from the American Academy of Periodontology In-Service Examination, spanning clinical photographs, histopathology, radiographs, cardiac rhythm strips, and anatomical illustrations. A sequential two-phase experimental design was used: in Phase 1, each model independently described each image, selected an answer, and provided a supporting citation; in Phase 2, applied only to questions answered incorrectly, models were given an expert-validated visual description and asked to re-answer, allowing diagnostic improvement through visual correction to be measured directly. Expert ground truth for image content was established by a board-certified periodontist and independently validated by a second board-certified periodontist. Model outputs were classified using a dual-process error taxonomy adapted from Norman's model of diagnostic reasoning, distinguishing perceptual errors, arising from inaccurate visual feature extraction, from cognitive errors, arising from flawed reasoning despite accurate perception, with cognitive errors further subdivided into correctable and persistent subtypes, and additional categories capturing compound perceptual-cognitive failures and compensatory reasoning that overcame inaccurate perception. Diagnostic accuracy and error type distribution varied significantly across models and image modality. Correcting inaccurate visual descriptions in Phase 2 improved diagnostic accuracy for a subset of previously incorrect responses, indicating that a meaningful share of errors originated at the level of visual perception rather than clinical reasoning; conversely, a distinct subset of errors persisted despite accurate corrected visual input, indicating reasoning-level failures independent of perceptual accuracy. Some models also reached correct answers despite generating inaccurate image descriptions, reflecting compensatory reasoning resilient to perceptual error. These findings show that aggregate accuracy scores conflate mechanistically distinct failure modes, and that perceptual and cognitive errors carry different implications for how MLLMs might be safely deployed or improved for diagnostic image interpretation. The expert-guided visual correction framework introduced here provides a generalizable, mechanism-based approach to benchmarking multimodal AI diagnostic performance that extends beyond periodontics to other visually driven diagnostic domains in medicine. As MLLMs become increasingly accessible to clinicians, residents, and dental educators, distinguishing perceptual from cognitive failure is essential for guiding responsible clinical use, targeting model refinement, and informing AI-augmented dental education and competency assessment.

13
Accuracy of a Smart-Ring VO2max Estimate and Five Published Prediction Equations Against Cardiopulmonary Exercise Testing: Development and Validation Study With Population-Scale Analysis

Dhawale, N.; Mukundan, S.; Agarwal, A.; Mondal, D.; Shanmugam, A.; Kumar, P.; Mittal, M.; Narasimhan, V.

2026-07-17 sports medicine 10.64898/2026.07.16.26358226 medRxiv
Top 0.1%
34.3%
Show abstract

Background. Maximal oxygen uptake (VO2max) is a leading marker of cardiorespiratory fitness and a strong predictor of all-cause mortality. Cardiopulmonary exercise testing (CPET) is the reference method but is resource-intensive, so consumer wearables estimate VO2max from passively collected signals; these estimates compress the fitness range, returning near-correct group averages while ranking individuals poorly. No peer-reviewed validation of a smart-ring VO2max estimate against CPET has been reported, and none in a South Asian cohort. Objective. To validate the Ultrahuman Ring AIR VO2max estimate against laboratory CPET, benchmark it against published prediction equations, and assess its generalization and construct validity. Methods. In a single-site paired ring-CPET cohort (N = 101; mean CPET peak VO2 43.3 mL{middle dot}kg-{superscript 1}{middle dot}min-{superscript 1}, SD 9.9), peak oxygen uptake was measured by treadmill or cycle-ergometer CPET, and the Ultrahuman Ring AIR estimate was computed from passively collected signals using a transparent ensemble based on published equations. Ensemble weights and calibration were selected on an 85-subject development set by an automated search minimizing a composite 5-fold cross-validated error criterion; the locked estimate was evaluated on a 16-subject held-out test set. The calibrated coefficients are proprietary. Agreement was quantified with mean absolute error (MAE), bias, Pearson r, regression slope and Lin's concordance correlation coefficient (CCC; bootstrap 95% CIs), and Bland-Altman limits of agreement. Separately, in 181,133 de-identified Ring users (no CPET reference), construct validity was assessed against ring-measured sleep, continuous glucose monitoring (n = 2,597), and a venous blood panel (n up to 15,203), adjusted for age, sex, and BMI, with lipoprotein(a) as a pre-specified negative control. Reporting followed TRIPOD and STARD. Results. With a self-reported fitness level provided, the estimate agreed with CPET peak VO2 at MAE 4.68 mL{middle dot}kg-{superscript 1}{middle dot}min-{superscript 1} (95% CI 3.93 to 5.49), Pearson r 0.79, CCC 0.79, and slope 0.71. The five published equations were worse on every metric (MAE 6.2 to 10.6, CCC 0.28 to 0.56, slope 0.32 to 0.42), each compressing the fitness range. On the held-out test set (n = 16), agreement held (r 0.84, slope 0.81, MAE essentially unchanged). Without the fitness input, full-cohort MAE was 5.16, still ahead of every published equation. At population scale, higher estimated fitness tracked a healthier profile on measurements the estimate does not use: better ring-measured sleep; higher continuous-glucose time in target range (79.6% versus 61.5%, top versus bottom decile; n = 222 and 399 of 2,597 users); and lower triglycerides, fasting glucose, and HOMA-IR (n up to 15,203 assayed per marker). These associations held after adjustment for age, sex, and BMI, whereas the pre-specified negative control lipoprotein(a) did not separate the deciles. Conclusions. The Ultrahuman Ring AIR VO2max estimate agreed with laboratory CPET substantially better than published prediction equations, held its agreement on held-out subjects, and ordered a large population along independent cardiometabolic gradients consistent with true fitness.

14
Vision and Language Models for Classifying Maxillary Sinus Disease on Cone-Beam Computed Tomography: A Transparent Multimodal Benchmark

Al-Hebshi, S.; Khalifa, H.; Pham, T. D.

2026-08-12 dentistry and oral medicine 10.64898/2026.08.11.26360189 medRxiv
Top 0.1%
32.3%
Show abstract

Background: Cone-beam computed tomography (CBCT) frequently captures the maxillary sinuses incidentally, and reliable automated detection of sinus abnormality is clinically relevant. Unlike most vision-language benchmarks in medical imaging, which pair images with pre-existing, human-authored clinical reports, findings text can also be generated directly by a large language model from the image itself--raising the question of how much diagnostic value such AI-derived text carries, and whether that value depends on independent verification. Multimodal artificial intelligence (AI) benchmarks risk overstating performance if the provenance of each input--image, raw AI-generated text, or radiologist-verified text--is not clearly separated and reported. Methods: We used 300 mid-sagittal CBCT slices from the MMDental dataset. ChatGPT generated findings text and a provisional normal/abnormal label for every slice (majority vote, three independent readings from the image alone); primary classification performance was assessed on this full, unfiltered set (n=300). A radiologist then independently reviewed each case's image together with ChatGPT's description, producing their own diagnosis; three cases were excluded as insufficient, yielding 297 confirmed cases. On this subset, every model was retrained and re-evaluated under identical 10-fold cross-validation on both the provisional ChatGPT-only labels ("pre") and the radiologist-confirmed labels ("post"), isolating the effect of label provenance from image or architecture. Eight vision architectures, seven language classifiers, and five VLMs were evaluated throughout; three generative models performed exploratory note-drafting. Findings: Raw ChatGPT-generated text produced the highest performance of any modality or condition: language models reached near-ceiling AUC (0.992 to 1.000, n=300), exceeding every vision model (AUC 0.799 to 0.880) and every VLM image-only probe (AUC 0.63 to 0.69). On the 297-case pre/post analysis, this advantage depended heavily on label source: language and text-derived VLM performance fell substantially from ChatGPT-only to radiologist-confirmed labels (e.g. BERT-base AUC 0.999 to 0.837), while vision-model performance was stable or modestly improved (e.g. DenseNet-121 0.867 to 0.891). The radiologist reclassified 62 of 297 cases (21%) relative to ChatGPT's provisional read, and a meaningful proportion of raw ChatGPT text was clinically uninterpretable or unsupported by the imaging. Interpretation: As shown here for the first time, raw, image-derived AI-generated text yields the highest apparent classification performance in this benchmark, but this reflects the text's alignment with its own self-generated labels rather than verified diagnostic content, and a substantial share of that text is not clinically explainable. Radiologist-confirmed text and labels give a lower but trustworthy estimate of true performance, on which convolutional neural network (CNN) vision models remain a stable, comparatively inexpensive baseline. Multimodal dental AI should report performance separately by modality and label provenance rather than pooling headline metrics.

15
Digital inclusion, access barriers and trust calibration in smartphone-based hypertension screening: a mixed-methods policy and implementation study in northern Nigeria

Dasa, D.; Davies, P.

2026-08-10 health informatics 10.64898/2026.08.07.26359947 medRxiv
Top 0.1%
31.4%
Show abstract

Objectives. To assess how digital inclusion factors and physical access barriers are associated with user trust in smartphone-based remote photoplethysmography (rPPG) hypertension screening, and to identify implications for digital health pol- icy, procurement and implementation in low-resource settings. Methods. Cross-sectional mixed-methods survey in five outpatient clinics in Kebbi State, northern Nigeria (N =287). Trust was measured using comfort, confidence and perceived usefulness Likert scales. Primary analyses used binary logistic models with HC3 robust standard errors; sensitivity analyses are reported in supplementary material. Free-text responses were thematically analysed. Results. Smartphone ownership was 51.2%; Transsion-brand devices comprised 56.5% of owners. Greater distance to a blood pressure facility was independently associated with lower perceived usefulness (OR 0.51, 95% CI 0.30-0.87; p=0.013) and lower comfort (OR 0.61, 0.37-0.98; p=0.042). Among owners, Transsion versus Samsung showed higher confidence odds (OR 3.82, 1.02-14.27; p=0.046). Qualitative themes supported the implementation interpretation: platform-fit and device speed requests among Transsion owners; connectivity and offline-first concerns among those with greater travel distance. No brand contrast achieved FDR-adjusted significance; brand findings are exploratory. Conclusions. Digital health policy and health technology assessment for smartphone-based screening should incorporate local device ecology, connectivity constraints, physical access burden and trust-calibration safeguards. Pre-implementation assessment of these factors is necessary for equitable and safe rPPG adoption in low-resource health systems.

16
Adaptation and Psychometric Validation of a Facility-Level Tool to Assess Telemedicine Readiness in Primary Care

Escobar-Agreda, S.; Villarreal-Zegarra, D.; Reategui-Rivera, C. M.; Paredes-Gonzales, Y.; Rojas-Mezarina, L.

2026-07-10 health informatics 10.64898/2026.07.01.26356790 medRxiv
Top 0.1%
30.7%
Show abstract

Background: Telehealth has expanded rapidly, yet its sustained use beyond emergency responses remains uneven. Facility and organizational conditions are modifiable determinants of healthcare intervention implementation, and readiness assessments are key steps in the early adoption and further uptake process. However, many telemedicine readiness assessment instruments lack psychometric evidence, limiting their value for benchmarking, prioritizing investments, and monitoring progress at scale. Objective: To adapt and psychometrically validate the Telemedicine Readiness Inventory at the Facility Level (TRI-F). Methods: We conducted a cross-sectional analysis using data from an online assessment of a large sample of primary care facilities (PCFs) in Peru, held between December 2023 and March 2024. The analytic sample included 774 PCFs, with one designated respondent per facility completing the survey. Internal structure was evaluated using exploratory factor analysis and confirmatory factor analysis. Internal consistency was estimated using Cronbachs alpha and McDonalds omega. Measurement invariance was evaluated across the facility complexity and the respondant time working at the PCF. Criterion validity was examined using Spearman correlations between the five domain scores (Organizational readiness, Processes, Digital environment, Human resources, and Regulatory issues) and proportion of telemedicine modality among overall outpatient encounters. Results: Confirmatory factor models showed adequate fit across domains, with CFI values ranging from 0.959 to 0.999, TLI from 0.952 to 0.997, RMSEA from 0.028 to 0.065, and SRMR from 0.016 to 0.057, for the five domains assessed. Internal consistency was acceptable to high across all domains ( = 0.75-0.87; {omega} = 0.76-0.88). Measurement invariance was supported across the facility category and time working at the PCF, with {Delta}CFI values below 0.010. Criterion validity analyses showed positive but small correlations between all five domains and proportional telemedicine use (rs= 0.15-0.23; p < 0.001). Conclusions: The adapted tool showed satisfactory structural validity, internal consistency, and measurement invariance for assessing telemedicine readiness in PCFs. The availability of percentile-based norms supports interpretation and benchmarking. The instrument can support implementation planning, monitoring, and prioritization of technical assistance in primary care settings.

17
Design tensions in a two-sided marketplace for reusable digital therapeutics software components: a qualitative interview study

Kowatsch, T.; Melamed, S.; Nissen, M.; Merz, Y.

2026-07-20 health informatics 10.64898/2026.07.17.26358332 medRxiv
Top 0.1%
30.5%
Show abstract

Objectives To identify stakeholder-perceived design tensions in a two-sided marketplace for reusable digital therapeutics (DTx) software components and to use these tensions to propose alternative marketplace concepts. Methods We conducted 24 semi-structured interviews with digital health researchers and professionals. Data were analysed using hybrid deductive-inductive codebook thematic analysis. The Magic Triangle provided the initial deductive structure. One researcher coded all transcripts; a second independently applied the developing codebook to five transcripts to refine definitions and consistency. Seventeen parent themes were synthesized into 12 design tensions, which informed three author-generated marketplace concepts. Results Participants described trade-offs concerning target users and host, component scope and customization, quality labels, verification, geographic scope, pricing, interoperability, platform launch, risks and market niche. The resulting concepts emphasized a regional startup ecosystem, a research-oriented hybrid marketplace or a global marketplace with stricter entry requirements. Discussion The concepts combine the tensions in different ways and highlight competing priorities in governance, openness, assurance, scalability and early platform growth. Conclusion Stakeholders identified recurring design choices for a DTx software-component marketplace. The concepts provide hypotheses for prototyping and evaluation; the study did not test technical feasibility, market demand, regulatory acceptability or effects on development cost or time.

18
Feature Selection with Quantum Annealing for Biomedical Machine Learning Applications

Dudgeon, S. N.; Lee, S. J.; Durant, T. J.; Nelson, B.; Young, H. P.; Ohno-Machado, L.; Taylor, R. A.; Schulz, W. L.

2026-07-06 health informatics 10.64898/2026.07.02.26357174 medRxiv
Top 0.1%
30.3%
Show abstract

Feature selection is a commonly used method in biomedical artificial intelligence and machine learning to identify a subset of high-quality variables that can be used to train downstream predictive models. It has been suggested that quantum feature selection (QFS), which takes advantage of the properties of quantum computers, may better identify variables that are correlated with the outcome while simultaneously reducing redundancy between selected variables. However, there are a limited number of studies evaluating their performance, particularly in real-world data sets. Here, we assess the performance of two QFS methods compared to random forest (RF) feature selection based on feature stability and the performance of a downstream classification algorithm when used to predict urinary tract infections in the emergency department from 211 original features extracted from the electronic health record. We found that a quantum binary quadratic model (BQM) and constrained quadratic model (CQM) had similar performance to RF feature selection (median F1 score of 0.60, 0.61, and 0.61 respectively) when 10 features were selected for an XGBoost classification model. The BQM and RF also had similar feature stability (0.91 and 0.94, respectively) while the CQM had lower stability (0.72). These findings show that QFS can be used with large, clinical data sets to identify features with high stability and predictive performance. As the capacity and quality of quantum computers continue to increase, these methods may offer additional benefits to classical feature selection methods.

19
When Algorithms Prescribe: A Cross-Sectional Study of Quality, Misinformation, and Engagement in Statin-Related Content on TikTok

Gharibyan, I.; Ahner, E.; Shao, R.; Sharma, D.; Navarsartian Tazehkand, T.; Diep, J.; Assoumou, B.

2026-06-08 health informatics 10.64898/2026.06.04.26354962 medRxiv
Top 0.1%
30.0%
Show abstract

Background: Statins are key to preventing atherosclerotic cardiovascular disease and lowering low-density lipoprotein cholesterol and cardiovascular events. However, skepticism regarding their safety and value persists and is increasingly influenced by social media. TikTok has emerged as a major source of health information, but its content varies in quality and accuracy. This study evaluated the quality, attitudes, misinformation, and engagement of statin-related content on TikTok. Methods: Public TikTok videos were collected using predefined search terms and coded by creator type, thematic content, and overall attitude. Video quality was assessed using the DISCERN instrument, the Patient Education Materials Assessment Tool for Audiovisual Materials, and the Global Quality Score. False or misleading claims were independently reviewed by two cardiology fellows. Associations between engagement and quality were also examined. Results: Of 1,349 screened videos, 258 met inclusion criteria. Most were educational (91.0%), with non-physician healthcare providers (34.5%) as the largest creator group. Risks or negative effects were discussed more often than benefits (63.2% vs 42.2%), and 39.5% contained at least one false or misleading claim, most often from complementary and alternative medicine providers and wellness promoters. Quality differed by creator type across all instruments, with physician-created content scoring highest. Video popularity showed minimal association with informational quality. Conclusion: Statin-related TikTok content frequently emphasizes harms, often contains misinformation, and varies substantially in quality by creator type. Greater involvement of healthcare professionals on social media may help improve digital health literacy and counter misleading information about statin therapy.

20
Automating the triage of rheumatology outpatient referrals: a comparative evaluation of 23 large language models under simple and advanced prompting

Roberts, L.

2026-08-10 health systems and quality improvement 10.64898/2026.08.05.26359488 medRxiv
Top 0.1%
27.4%
Show abstract

Objective. Triage of rheumatology outpatient referrals is a high-volume administrative task that consumes senior specialist time without advancing patient care. The human triage system is only moderately accurate and reproducible. We assessed whether contemporary large language models (LLMs) are able to perform well enough to support automating this task in practice. In addition, the effects of different prompting techniques on triage accuracy and cost was assessed to help identify to optimal approach. Methods. Twenty referral scenarios spanning the urgency spectrum, based on real referrals were created by a certified Australian rheumatologist. Four rheumatologists triaged all cases independently and blinded, to produce a consensus reference standard. Twenty-three LLMs each triaged every referral into one of five urgency categories, three times (1380 outputs per condition). The experiment was run with a simple prompt and repeated with a advanced prompt supplying explicit triage expectations and worked examples. Results. All 2760 attempts returned valid categories. Under the simple prompt, performance separated into distinct tiers, larger models were more accurate (Spearman rho=0.42; P=.047) and accuracy tracked cost. Advanced prompting minimised between-model variance in accuracy 5.3-fold (0.014 to 0.003; Levene P=.01), abolished the size-accuracy association (rho=-0.05; P=.83) and removed the accuracy-cost relationship. Leading models matched expert consensus on most cases, within or above the range reported for human triage. Under-triage errors persisted with some LLMs. Conclusion. Contemporary LLMs categorise rheumatology referral urgency as well or better than published human triage systems. Advanced LLM prompting methods substitute for the reasoning capability of larger models, suggesting that LLM performance on this task may not require the most expensive models. The tools to automate this administrative task appear to already exist. Strong candidate LLMs that might serve a production ready solution have been identified.